Journal of Mathematical Biology
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Journal of Mathematical Biology's content profile, based on 40 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Katsaounis, D.; Chaplain, M. A.; Sfakianakis, N.
Show abstract
Cancer progression is driven by the interplay between cancer cell phenotypic variability regulated by EMT/MET, and cell-cell and cell-matrix interactions. Starting with an individual-based model, in which every cell is characterised by its position, velocity and a continuously varying epithelial-mesenchymal phenotype, we derive, via a kinetic description and a mean-field limit, two alternative macroscopic formulations: an Euler-like system that couples mass and momentum to the phenotypic variable, and a single advection-aggregation-diffusion equation (AADE) for the cancer cell density. Both macroscopic models retain the non-local adhesion-repulsion forces, haptotactic response to an evolving ECM, and phenotype-dependent transition dynamics driven by TGF-{beta}. Numerical experiments in two spatial dimensions indicate that the macroscopic equations reproduce key scenarios obtained at the individual scale. In particular, a microscopic-macroscopic comparison shows that the AADE density reproduces acurately both the spatial localisation and the phenotypic decomposition of the individual cell population. We also demonstrate that varying only the steepness of the TGF-{beta} switch function, changes the EMT response from an almost binary epithelial-mesenchymal (EM) separation to a partial EM phenotypes. This study provides a systematic bridge from stochastic, heterogeneous cell dynamics to continuum descriptions, for investigating phenotype driven tumour invasion and supporting the choice of macroscopic models in large-scale simulations and analytical studies.
Ma, S.; Li, Y.
Show abstract
The global potential of a chemical reaction network has many applications and is closely related to the stochastic detailed balance. However, many fundamental questions concerning stochastic detailed balance remain unresolved, such as whether it depends on the system volume and how to construct new systems that satisfy it. In this paper, we show that stochastic detailed balance may depend on the system volume. We therefore introduce four types of stochastic detailed balance according to their dependence on volume and rate constants, and systematically investigate the relationships among them. Our results distinguish detailed balance arising from particular choices of volume and parameters from that enforced by network structure, and identify conditions under which detailed balance at one volume extends to all volumes. We further obtain a class of networks satisfying stochastic detailed balance for every volume and every positive choice of rate constants, and construct new systems whose global potentials exhibit double-well structures.
BV, H.; Adigwe, S.; Jolly, M. K.; Gedeon, T.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCell fate decisions are driven by gene regulatory networks (GRNs). While the mutually inhibitory toggle switch effectively models binary fate decisions, fully connected inhibitory networks with more than two nodes fail to capture multi-fate decisions due to the low prevalence of "single high states", where only a single master regulator is highly expressed. The goal of this study is to find network structures that support all single high states. We find that the only network that attains the highest possible prevalence of all single high states within the set of monotone Boolean (MB) models is completely disconnected. Since biological networks typically require connectivity, we investigate network structures that support equipotency, where all single high states have equal prevalence within MB models. Finally, we characterize the networks that support multistability between all single high states, finding that it is possible only in networks in which each node either has self-activations or is inhibited by every other network node. Our findings provide a theoretical framework for understanding the network design principles that can support simultaneous differentiation into multiple distinct cell types.
Huang, Q.; Guo, H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCellular automata and graph reaction-diffusion systems encode local spatial interactions in different mathematical forms. We develop a cochain-operator calculus for these two settings. Over a finite field Fq, every local rule on a finite neighborhood has a unique reduced polynomial representative. On an oriented line, the coboundary and endpoint maps recover the left and right shifts. Our main theorem shows that these operators, together with linear operations, constant cochains, and the degree-zero cup product, generate every finite-radius polynomial cellular automaton. Explicit formulas for Rules 30, 110, and 22 show how reflection-invariant linear coupling, directed transport, and nonlinear neighbor interactions enter the calculus. On a general graph, d*d is the unweighted combinatorial Laplacian and enters a graph reaction- diffusion recurrence. Over [R], the term - Dd*d with D [≥] 0 admits the usual diffusion interpretation; over Fq, the corresponding expression defines modular coupling without an intrinsic order. In the morphogenetic examples, we therefore distinguish pattern-generating dynamics from finite-state observation and use the Betti numbers of active induced subcomplexes to summarize observed patterns. This yields a common algebraic representation without identifying real-valued diffusion with finite-field dynamics.
Sadhukhan, S.; Santra, D.
Show abstract
Diffuse gliomas are deadly because the individual tumor cells invade - they travel far from the imageable mass, so it is impossible to remove the tumor completely. On the cellular level, glioma cells seem to be in either a "go" state (in which they do not divide) or a "grow" state (in which they do not migrate). We investigate what this tiny choice has to say about the large-scale speed of the invasion front and whether the implication is sufficiently strong to rule out the classical description of the Fisher-Kolmogorov-Petrovsky-Piskunov (Fisher-KPP) type, in which a single phenotype migrates and proliferates. We derive a two-phenotype reaction-diffusion model with density-dependent switching, and we prove the cooperative (quasi-monotone) structure and the associated comparison principle and study travelling-wave solutions of the model. A leading-edge linearization gives minimal front speed as minimizer of an explicit dispersion relation, and direct simulation verifies the predicted speed. In the experimentally relevant fast switching limit, we find a closed-form expression for the speed, that is, we obtain an effective Fisher-KPP equation with rescaled diffusivity and growth rate, with the fractions of the phenotypes. The "go-or-grow" (GoG) front can move at a maximum speed of half the Fisher speed for the same single-cell motility $D$ and proliferation rate $r$, which occurs only when the cells divide their time equally between the two phenotypes. This bound is directly testable: measurement of the front speed, plus independent determination of $D$ and $r$, discriminates the two hypotheses, and in the GoG case, yields recovery of the phenotype balance. We then extend the result to anisotropic (DTI-informed) invasion along white-matter tracts and discuss implications for understanding clinical measurements of growth rate.
Pizarro Galleguillos, F.; Bhonsale, S.; VAN IMPE, J.
Show abstract
The dynamics of gene regulatory networks are governed by intrinsic noise, stemming from the random nature of biochemical reactions, and by extrinsic noise, arising from fluctuations in cellular components and environmental conditions. Together, these sources can compromise the reliability of predictive computational models if not properly accounted for, and capturing both effects within a single framework remains a non-trivial task in computational biology. In this work, we propose an uncertainty quantification framework that addresses these two contributions jointly: intrinsic stochasticity is described through a partial integro-differential equation (PIDE) for the protein probability density function, whereas extrinsic noise is represented as parametric uncertainty in the kinetic parameters. The propagation of the uncertainty is carried out via an intrusive polynomial chaos expansion (PCE), in which the PCE coefficients are obtained from a stochastic Galerkin projection of the PIDE, yielding a coupled deterministic system that is solved with standard numerical methods. We illustrate the approach on a positive autoregulatory gene network with one and two uncertain kinetic parameters. The proposed approach accurately reproduces the mean, variance, and full protein probability density function, including the bimodal distributions, at a substantially lower computational cost.
van de Beek, H.; Beldjenna, M.; Fidler, M. L.; Zwep, L. B.; van Hasselt, J. G. C.
Show abstract
Asymptotic standard errors for the parameters of a nonlinear mixed-effects model fitted by first-order conditional estimation (FOCE) or FOCE with interaction (FOCEI) require the observed (Fisher) information -- the negative second derivative of the population objective at the optimum. The gradient of this objective can be computed exactly from sensitivity equations, but the observed information is conventionally still formed by finite differencing, which is less accurate and step-size dependent. Our objectives are to (i) derive the FOCE and FOCEI observed information in closed form within the same sensitivity-equation framework, and (ii) quantify the precision this recovers. Writing the objective as a data term plus the log-determinant of the first-order inner Hessian, the population Hessian splits so that the data term reuses the second-order sensitivities already needed for the gradient, whereas the log-determinant term requires third-order sensitivity equations -- confining the third-order dependence to a single term, where it enters in exactly two places. A finite-difference error analysis shows the differenced Hessian attains an accuracy no better than the square root of the objectives evaluation accuracy, whereas the analytic form is limited only by the sensitivity and differential-equation solutions, with no step size to tune. We illustrate this on a one-compartment oral model with first-order absorption fitted to warfarin data, where the differenced standard error is usable only over a narrow band of step sizes while the analytic value carries none. Implemented in the open-source R package nlmixr2, the method makes exact, reproducible standard errors routine, supporting more dependable confidence intervals, identifiability assessment, and uncertainty propagation.
Masutomi, Y.;Kobayashi, K.
Show abstract
The photosynthesis-transpiration-stomatal conductance (An-E-gs) model framework is widely used for estimating photosynthesis, transpiration, and stomatal conductance in plants. The model equations are solved by numerical iteration, and the converged model values are deemed the solution. However, there has been no general guarantee that the iterative procedure converges to a solution or that the procedure leads to convergence. Building on the recent proof of the existence of a unique set of solutions, we herewith propose a numerical algorithm that is guaranteed to converge to the solution for the An-E-gs model framework. We first analytically prove that the proposed algorithm necessarily converges to a solution. We then demonstrate the convergence across contrasting combinations of leaf temperature, relative humidity, light, atmospheric CO2, and wind speed. We further demonstrate rapid convergence with the algorithm: no more than ca. 10 iterations for approximately 10-3 mol CO2 m-2 s-1 precision in net photosynthesis and no more than ca. 20 iterations for 10-7 mol CO2 m-2 s-1 precision. By guaranteeing convergence to the solution, this algorithm eliminates concerns about nonconvergence in leaf gas-exchange calculations and is expected to serve as a robust foundation for a range of studies from leaf-level gas exchange to global-scale carbon and water cycle dynamics.
Schreiber, S.; Brennan, J.; Spaak, J. W.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_LICommunity assembly graphs (CAGs) summarize which species combinations can coexist and how single-species invasions drive transitions between them, encoding the pathways, alternative endpoints, and cycles that make up a communitys assembly history. Constructing CAGs from dynamical models requires methods that are both computationally tractable and faithful to the underlying ecological dynamics. However, existing methods rely on restrictive assumptions, such as global stability, that exclude alternative stable states and non-equilibrium dynamics known to occur in empirical systems. C_LIO_LIWe develop a computational pipeline that constructs CAGs from any generalized Lotka-Volterra model. Building on the invasion graph framework and its connection to permanence, the pipeline verifies that community dynamics are bounded, identifies which subsets of species coexist in the sense of permanence, determines which single-species invasions are dynamically realized, and assigns each community a topographic height equal to the length of the longest assembly path leading to it. We also provide a numerical algorithm to simulate the dynamics of community assembly. C_LIO_LIWe prove several general properties of the resulting graphs, including that a successful invader is never subsequently excluded and that, in the absence of assembly cycles, permanent communities can be reassembled by introducing their species one at a time in the right order. We prove that the CAG faithfully reproduces the compositional shifts seen in the numerically simulated dynamics of assembly. Applying the pipeline to three empirically based models (a New Zealand grassland, a European pasture, and a Puerto Rican ant community), we show how competition strength and mutualistic feedbacks reshape the assembly landscape and how intransitive competition generates assembly cycles. C_LIO_LIOur approach accommodates alternative stable states and non-equilibrium dynamics without requiring global stability, and it turns the long-standing landscape metaphor into a quantitative, mechanistically grounded object by resolving what "height" means. More broadly, it makes the topography of the assembly pathways measurable, providing a way to compare the historical contingency and predictability of the assembly in ecological systems. C_LI
van Boven, M.; Bootsma, M. C.
Show abstract
Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.
Kringelbach, M. L.; Deco, G.
Show abstract
Brain dynamics can be described in three different convenient mathematical languages, namely connectome harmonics, turbulence and complex harmonics (CHARM). Here we demonstrate that these theoretical frameworks can be rigorously unified, under the functional calculus, as one self-adjoint operator and its single spectral measure. The connectome Laplacian carries that measure; the harmonics are its spectral projections, the turbulence smoothing kernel is its resolvent, and the CHARM form is its unitary propagator. The bridge that makes this exact is a textbook fact: The exponential distance rule, which is the empirical kernel of the turbulence model, is the Greens function of a screened Laplacian, so the local order parameter is the phase field passed through the resolvent of the same operator whose eigenfunctions are the harmonics. A single shared control parameter, the spectral gap, simultaneously yields the cortical hierarchy, the turbulent information cascade and the structured interference the CHARM form measures. This unification makes a strong predictive claim. If the harmonic projections, the turbulence resolvent and the CHARM propagator really are three functions of one operator, then any structural perturbation that re-tunes the operator must move all three signatures in unison and must do so with a single coupling. We test this prediction with a pharmacological perturbation by lysergic acid diethylamide (LSD), which is known to change the emotional state, by empirically perturbing the operator with a 5-HT2A receptor density map and asking whether one scalar coupling can simultaneously predict the multi-scale turbulence shift observed, through the resolvent, and the macroscale harmonic energy redistribution, through first-order Rayleigh-Schrodinger perturbation theory. We found that the two independent functional domains respond in unison to one structural perturbation of one operator. The identity is exact as operator calculus and its purchase on the brain depends on a single load-bearing seam, the degree heterogeneity of the connectome, which we make explicit. We propose that this single-operator structure is the necessary mathematical scaffolding of our Entangled Loop theory.
Brinas-Pascual, N.; Alarcon, T.; Calvo, J.; Guerrero, P.; Oliver-Bonafoux, R.
Show abstract
The study of tissue dynamics has been stimulated during the last decades thanks to the use of quantitative descriptions, with the development of several theoretical and computational frameworks, many of them revolving around the notion of reaction-diffusion systems, eventually with additional structure variables beyond time and space. The use of structure variables can accommodate phenotypic traits. In this work, we study a family of competition models, where a given population depends on a resource (e.g. oxygen) and several populations are competing for it. Our quantitative description incorporates phenotypic traits and heterogeneity at the level of cell cycle variations, which influence replication rates via oxygen consumption. This enables us to replicate the fitness of specific subpopulations to environmental conditions (e.g. oxygen shortage or external influences). Using numerical simulations, we show that such models display dynamical pattern formation in the form of coupled travelling wave profiles that expand or retreat at the same wave speed. The full theoretical analysis of such dynamics is quite involved; to circumvent this difficulty, we introduce a quasi-stationary approximation for the resource dynamics. We find that this approximation can reproduce the overall behaviour very accurately, with the additional benefit of allowing theoretical treatment of the reduced model. In this way, we provide estimates on the wave speed which are numerically shown to be robust across a wide range of macroscopic parameters of the full model. The wave speeds are thus found to depend strongly on the proliferation rate of the fittest population, resembling a winner-takes-all dynamics.
Nayeem, J.; Salek, M. A.; Biswas, M. H. A.; Kabir, M. H.
Show abstract
Background: Tuberculosis remains a persistent infectious disease whose control is complicated by latent infection, delayed treatment, incomplete recovery, reinfection, and continuing transmission from infectious individuals. Although treatment is central to tuberculosis management, it is frequently represented only as a transition parameter in mathematical models rather than as a separate epidemiological state. In this study, treatment was therefore incorporated explicitly as an independent compartment so that its influence on transmission, recovery, disease-induced mortality, and long-term disease persistence could be evaluated. Methods: A deterministic nonlinear compartmental model was formulated by dividing the total population into susceptible, exposed, actively infected, treated, and recovered classes. Reinfection of recovered individuals, progression from latent infection to active disease, movement of infectious individuals into treatment, treatment-associated recovery, natural mortality, and disease-induced mortality were included. Positivity and boundedness of the solutions were examined to establish biological validity. The basic reproduction number, R0, was derived through the next-generation matrix approach. Disease-free and endemic equilibria were determined, and their local and conditional global stability properties were investigated using Jacobian analysis, the Routh-Hurwitz criterion, center manifold theory, Lyapunov functions, and LaSalles invariance principle. Normalized sensitivity indices, Latin hypercube sampling, partial rank correlation coefficients, and numerical simulations were also applied. Results: The disease-free equilibrium was shown to be locally asymptotically stable when ,R0<1 whereas sustained transmission and a unique endemic equilibrium were associated with R0>1. Under the stated reduced-model assumptions, stability of the endemic equilibrium was established. Transmission-related parameters were identified as the strongest positive contributors to disease persistence. In contrast, treatment and recovery parameters were found to reduce the reproduction number and infectious burden. Numerical simulations indicated that stronger treatment implementation and reduced transmission opportunities produced substantial reductions in active tuberculosis cases. Conclusion: Treatment was shown to function as both a clinical pathway and an epidemiological control mechanism. The proposed framework may support the design of treatment-centered strategies for reducing tuberculosis prevalence and preventing long-term endemic persistence.
Ergon, R.
Show abstract
The general random walk model (GRW) of Hunt (2006) is used to infer directional evolution in mean trait values from sparse fossil data by modeling phenotypic change as the accumulated result of small steps with mean step sizes and step variances. Using simulations and real data cases, Ergon (2026) showed that the step variances can be estimated reasonably well only when the mean trait values have small measurement errors, while for fossil data with realistic measurement errors they appear to be extremely difficult to find, and they are often found to be negative. In the simulations Ergon (2026) assumed that the true phenotypic mean values were known. Here, I essentially repeat these simulations under the assumption that only mean trait values with large measurement errors are known, and based on weighted mean squared error (WMSE) comparisons the conclusion is that weighted least squares (WLS) is a better method than GRW. A second conclusion is that WLS is a better method also in the possibly rare cases with large measurement errors where the GRW parameters are estimated well. The GRW method is simply not flexible enough to handle such cases. A third conclusion is that Akaike Information Criterion (AIC) results for GRW models with large measurement errors relative to the step variance may be overly optimistic.
Gutierrez, M. A.; Gog, J. R.
Show abstract
In a population model for an infectious disease, we consider the early stochastic dynamics of an emergent 'mutant' strain, appearing and spreading during an epidemic of another 'wildtype' strain. The mutant may not reach establishment in the host population. The time at which the mutant first appears determines its probability of establishment. We calculate this establishment probability with two methods. The first method assumes a classical branching process, with a constant transmission rate. The second method reflects the changing size of the pool of susceptible hosts, due to the dynamics of the wildtype. We find that susceptible depletion can substantially impact the establishment probability. We explore the consequences of this stochastic establishment on the "escape pressure" acting on a pathogen to produce immune escape variants. We find that the overall escape pressure rate depends strongly on the appearance time of the mutant, especially if the establishment probability is itself shaped by the continued spread of the wildtype. In most scenarios, the escape pressure rate (and thus, the risk of new escape variants) peaks slightly earlier than the prevalence of the wildtype strain. Integrating the escape pressure over time, we obtain the cumulative escape pressure generated by the wildtype epidemic. The relationship between the escape pressure and the vaccination coverage depends on the cross-immunity, due to susceptible depletion. For example, with intermediate cross-immunity, the risk of immune escape may be lowest at intermediate vaccination coverages. Thus, these results raise important considerations for vaccination strategies in response to novel outbreaks.
Oosawa, C.
Show abstract
Biochemical Systems Theory (BST) represents nonlinear biochemical rate laws by local power-law approximations in logarithmic concentration coordinates. First-order coefficients are elasticities, whereas higher-order derivatives describe local log-synergism and its variation. Ordinary higher derivatives, however, are not tensorial under nonlinear reparameterizations and can mix biochemical response structure with coordinate artifacts. We formulate a covariant hierarchy, cBST1-cBST3, on a positive operating-point space equipped with a declared reference connection. cBST1 recovers classical elasticities in a flat logarithmic chart, cBST2 is the covariant Hessian of the log-response, and cBST3 is the symmetrized covariant derivative of cBST2. The framework is quantitatively evaluated using three representative rate laws from the curated yeast glycolysis model BIOMD0000000064: glucose transport, glucose phosphorylation, and phosphofructokinase. For 10,000 finite log-concentration perturbations at each of five radii, cBST2 reduced the cBST1 log-rate root-mean-square error by 96.6-98.6% at the largest tested radius, and cBST3 provided a further 68.4-98.1% reduction. The contracted cBST2 and cBST3 terms strongly predicted the corresponding lower-order truncation errors. Under the nonlinear transformation qi = sinh(ui), covariant contractions agreed across coordinates to within a 95th-percentile relative error of 1.2 x 10-13, whereas ordinary higher derivatives showed order-unity coordinate mismatches. Supplementary analytic tests recovered the expected second-, third-, and fourth-order truncation-error scaling. These results show that cBST1-cBST3 are not only coordinate-consistent descriptors but also practical diagnostics of where local power-law approximations require higher-order correction.
Asgedom, A.;Kefela, Y.
Show abstract
Cancer remains a global health challenge requiring sophisticated understanding of tumor-immune dynamics for effective treatment design. Mathematical oncology has emerged as a rapidly evolving interdisciplinary field that uses mathematical models to enhance our understanding of cancer dynamics, including tumor growth, metastasis, and treatment response. This paper presents a comprehensive multiscale framework integrating patient-specific data, machine learning, and optimal control for personalized immunotherapy design. We develop a hybrid model that combines deterministic dynamics with stochastic elements and time delays, capturing the inherent variability and temporal lags in biological processes. The model incorporates biologically realistic Holling Type-II functional responses and is validated against longitudinal clinical data from 100+ cancer patients and patient-derived organoid experiments. Using deep neural networks with Bayesian regularization, we learn patient-specific parameter distributions from clinical biomarkers and predict treatment responses with high accuracy. Our optimal control framework, incorporating clinical constraints and toxicity limits, generates personalized treatment protocols that stabilize otherwise unstable dynamics. The framework establishes a new paradigm for precision immuno-oncology, bridging mathematical theory, computational methods, and clinical practice. Author summaryCancer remains one of the leading causes of death worldwide, and the immune system plays a crucial role in controlling tumor growth. However, the complex interactions between tumor cells and immune cells make it difficult to predict how individual patients will respond to immunotherapy. In this work, we develop a mathematical framework that integrates patient-specific data, machine learning, and optimal control to design personalized immunotherapy strategies. Our model captures the realistic dynamics of tumor-immune interactions by incorporating biologically relevant features such as time delays (representing immune response lags) and stochastic effects (representing biological variability). Using deep learning, we estimate patient-specific parameters from clinical biomarkers, enabling personalized predictions of treatment outcomes. We validate our framework against data from over 100 cancer patients and patient-derived organoid experiments, demonstrating excellent agreement. Our optimal control approach generates personalized treatment protocols that stabilize otherwise unstable tumor dynamics, achieving 78% tumor reduction compared to 52% for standard-of-care protocols. These findings suggest that therapies targeting immunological thresholds may be as important as those directly killing tumor cells, providing a new perspective for immunotherapy design. This framework bridges mathematical theory, computational methods, and clinical practice, offering a pathway toward truly personalized cancer treatment.
Rajakumar, A.; Buenzli, P. R.; Simpson, M. J.
Show abstract
Understanding and predicting extinction risk is a central challenge in population biology. Mathematical models incorporating Allee thresholds are commonly used to understand population dynamics and to assess extinction risks. Inaccurate predictions can have serious consequences for conservation management. In this simulation study, we develop a likelihood-based inference and prediction workflow to estimate parameters, including the Allee threshold and population diffusivity parameters, using noisy count data generated using a well-defined discrete model. Although parameters are identifiable according to commonly used criteria, the accuracy of resulting predictions depends strongly on the quantity, quality, collection time and spatial resolution of the data. Our workflow demonstrates that seemingly reliable parameter estimates can lead to inaccurate predictions, highlighting the need for careful consideration of data quality and quantity to guide extinction-risk modelling and prediction. Open source software is provided on GitHub to replicate and extend all results considered.
Sturrock, M.; Shahrezaei, V.
Show abstract
Approximate Bayesian computation sequential Monte Carlo (ABC-SMC) propagates its particles with a perturbation kernel, and with the standard Normal kernel it degrades sharply as the parameter dimension grows, a failure usually attributed to dimension itself. We show instead that it is governed by the quality of the summary statistics, with dimension entering only through a separate and milder mechanism, and that the two must act together for the Normal kernel to break. The first ingredient is covariance overinflation: the kernel covariance, estimated from the particle cloud, overshoots the true posterior covariance by a factor set by information loss in the summary statistics. We derive this overscaling factor in closed form for a Gaussian model with sufficient statistics and show that it stays modest at any dimension, shrinking toward its baseline value as the tolerance tightens; the extreme values seen in practice (of order 103) are a signature of insufficient summaries, not of dimension. The second ingredient is perturbation overconcentration: the normalised Normal step size concentrates around one as the dimension grows, so every proposal overshoots by the same factor. Either ingredient alone is harmless; only their combination breaks the Normal kernel. A Cauchy kernel (multivariate t with one degree of freedom) removes the concentration, keeping a positive acceptance rate under arbitrary overscaling at a bounded worst-case cost of 1.87x in expected squared jump distance. In a Metropolis-Hastings framework we derive closed-form acceptance rates for both kernels that illustrate the advantage of the Cauchy kernel in this limit. A series of full ABC-SMC computational experiments on five problems at d = 12, including a hierarchical gene-expression model, show the Cauchy reducing the sliced Wasserstein distance to the reference posterior by factors of up to 50 with the same simulation budget. Since the summary statistics are commonly insufficient for the models that require ABC, overinflation is structural and the Cauchy perturbation kernel is the right default for problems in higher dimensions.
Wu, J.
Show abstract
Many comparative analyses operate on rectangular matrices whose columns represent the same variables across observations. Phylogenomic measurements, however, are attached to tree branches. Converting locus-specific trees into a common locus-by-coordinate matrix is straightforward only when loci contain the same taxa and compatible topologies. In real datasets, missing taxa can delete reference branches or collapse adjacent branches into composite coordinates, while gene-tree discordance can cause a coordinate that survives taxon restriction to be absent from the empirical tree. Without a formal account of these changes, non-equivalent quantities may enter the same column and distinct causes of missingness may be conflated. Using standard tree-restriction operations, I first construct a provenance-retaining coordinate ledger. For each locus, every reference edge is recorded as deleted, retained individually, or incorporated into a composite coordinate whose original-edge membership is preserved. Each surviving coordinate is then assigned a recovery state by asking whether its corresponding split is displayed in the normalized empirical tree. Composite member sets generated across loci define a common column set, yielding one locus-by-coordinate state matrix. The matrix is unique because the graph reduction is unique. Retained taxa span one minimal subtree; its degree-two vertices lie on uniquely determined nonbranching paths, so suppressing them in any order yields the same reduced tree and grouping of original edges. Each edge therefore has one fate, each composite coordinate one member set, and each empirical split query one answer. Thus every locus-by-coordinate cell has one determined state under fixed labelled inputs and conventions. Fiber-split equivalence shows that projected splits faithfully encode these graph-derived edge groups, explaining why split-based implementations recover the same coordinate structure and cell states. A deterministic exhaustive validation on the six-taxon worked-example domain evaluated 14,160 primitive and composite cells across 472 normalized empirical-tree cases. Separately implemented graph- and split-side evaluators agreed in every tested cell, and two clean runs produced byte-identical canonical result files. No branch lengths are required. When valid lengths are supplied, numerical entries may be added separately by single-edge lookup or a declared selected-component-edge sum. The theorem guarantees a unique coordinate-and-state matrix under the stated inputs, not an input-independent numerical matrix or historical truth. It provides an auditable foundation for branch-wise comparative analyses under heterogeneous taxon coverage and gene-tree discordance.